Papers with Wikipedia edits
Generating Inflectional Errors for Grammatical Error Correction in Hindi (2020.aacl-srw)
Copied to clipboard
| Challenge: | Automated grammatical error correction is a data-heavy task . indic languages have a relatively low amount of digitized content and complex morphology . |
| Approach: | They generate a corpus of inflectional errors for training neural networks to correct grammatical errors in Hindi. |
| Outcome: | The proposed model trains on a corpus of inflectional errors extracted from Wikipedia edits. |
WIKIBIAS: Detecting Multi-Span Subjective Biases in Language (2021.findings-emnlp)
Copied to clipboard
| Challenge: | a particular type of bias is subjective bias, which introduces improper attitudes or presents a statement with the presupposition of truth. |
| Approach: | They propose to annotate a Wikipedia edits corpus with 4,000 sentence pairs to detect subjective bias. |
| Outcome: | The proposed dataset can be used as a research benchmark and generalize to multiple domains. |
The Million Authors Corpus: A Cross-Lingual and Cross-Domain Wikipedia Dataset for Authorship Verification (2025.findings-acl)
Copied to clipboard
| Challenge: | Authorship verification (AV) is a crucial task for identity verification, accountlinking, historical linguistics, and AI-generated text identification. |
| Approach: | They propose to use Wikipedia's Million Authors Corpus to examine authorship verification models on a broad scale. |
| Outcome: | The proposed dataset includes 60.08M textual chunks, contributed by 1.29M Wikipedia authors. |